Papers with neural network models
Copied to clipboard
| Challenge: | a tutorial aims to introduce the nascent field of interpretability and analysis of neural networks in NLP . |
| Approach: | This tutorial will introduce the nascent field of interpretability and analysis of neural networks in NLP. |
| Outcome: | This tutorial will introduce the nascent field of interpretability and analysis of neural networks in NLP. |
Copied to clipboard
| Challenge: | Large pre-trained language models have hundreds of millions of parameters and take several gigabytes of memory to train and inference. |
| Approach: | They propose an open-source knowledge distillation toolkit designed for natural language processing that provides a set of predefined distillation methods and can be extended with custom code. |
| Outcome: | The proposed method is comparable with or even higher than the public distilled BERT models with similar numbers of parameters. |
Copied to clipboard
| Challenge: | AutoNLU is an on-demand cloud-based system that enables users to create and edit datasets and train and test different state-of-the-art NLU models. |
| Approach: | They introduce an on-demand cloud-based system that provides an easy-to-use interface . they build powerful keyphrase extraction models that achieve state-of-the-art results . |
| Outcome: | The proposed model achieves state-of-the-art on two public benchmarks and is easy to use and use. |
Copied to clipboard
| Challenge: | Unsupervised and semi-supervised learning has been addressed scarcely in NLP . this tutorial provides an introduction to variational inference followed by an example-driven discussion of how to use variational methods for training DGMs. |
| Approach: | This tutorial provides an introduction to variational inference followed by an example-driven discussion of how to use variational methods for training DGMs. |
| Outcome: | This tutorial provides an introduction to variational inference followed by an example-driven discussion of how to use variational methods for training DGMs. |
Copied to clipboard
| Challenge: | Recent years have seen a paradigm shift in neural text generation due to advances in deep contextual language modeling and transfer learning. |
| Approach: | They will discuss how and why NLG models succeed/fail at generating coherent text. |
| Outcome: | This paper will discuss how and why these models succeed/fail at generating coherent text, and provide insights on several applications. |
Copied to clipboard
| Challenge: | Existing approaches to use word embeddings for text generation have been limited. |
| Approach: | They propose to use GANs with word embeddings to reproduce writing style in text . they use a sentence embeddable vector to model people's way of expression . |
| Outcome: | The proposed model outperforms baseline text generation networks across several metrics including BLEU-n, METEOR and ROUGE. |
Copied to clipboard
| Challenge: | Inductive biases can arise from any aspect of the model architecture, study finds . we investigate which architectural factors affect how models generalize . |
| Approach: | They investigate which architectural factors affect generalization behavior of neural network models . they use English question formation and English tense reinflection as test cases . |
| Outcome: | The findings suggest that human-like generalization requires architectural syntactic structure. |
Copied to clipboard
| Challenge: | Existing deep neural network models lack mechanisms to highlight important sentiment terms. |
| Approach: | They propose a method to incorporate affective knowledge into deep neural network models by mapping affective influence vectors to an affective impact value and integrating them into long-term memory models to highlight affective terms. |
| Outcome: | The proposed approach improves on three large datasets by 1.0% to 1.5% on the benchmark datasets. |
Copied to clipboard
| Challenge: | a dataset of 17,000 manually labeled documents is large for determining entity-oriented polarity in business news. |
| Approach: | They propose a convolutional neural network-based approach to classify entity-oriented polarity in business news. |
| Outcome: | The proposed model is based on convolutional neural networks and is small on the scale of existing models. |
Copied to clipboard
| Challenge: | Existing studies on the effectiveness of attention in NLP do not consider changes in semantic capability of different components. |
| Approach: | They propose a framework that exploits a convex hull representation of sequence semantics in an n-dimensional Semantic Euclidean Space and defines indicators to capture the impact of attention on sequence semantic. |
| Outcome: | The proposed framework exploits a convex hull representation of sequence semantics in an n-dimensional Semantic Euclidean Space and defines indicators to capture the impact of attention on sequence semantic. |
Copied to clipboard
| Challenge: | Unsupervised parsing is a task that can be learned without substantial prior knowledge. |
| Approach: | They train an unsupervised model for Arabic, Chinese, English, and German to learn syntactic structure from unlabeled text. |
| Outcome: | The PRPN architecture outperforms trivial baselines and acquires at least some parsing ability for all languages. |
Copied to clipboard
| Challenge: | Neural networks are vulnerable to adversarial examples due to slightly perturbed input data. |
| Approach: | They propose a method that evaluates the robustness of text classification models by an optimization problem that identifies a minimum synonym swap that changes the classification result. |
| Outcome: | The proposed method achieves high scores in human evaluations of grammatical correctness and semantic similarity for an IMDb dataset and implements adversarial training with the IMD and SST2 datasets. |
Copied to clipboard
| Challenge: | Existing datasets for low-resource language Vietnamese assess semantic similarity . a dataset for word pairs with similarity levels is needed to evaluate these models . |
| Approach: | They present two new datasets for the low-resource language Vietnamese to assess models of semantic similarity. |
| Outcome: | The two datasets are comparable to the English datasets. |
Copied to clipboard
| Challenge: | Argument(ation) mining is a task of identifying argument structure from text . lack of training data makes it difficult to train models based on limited data sets. |
| Approach: | They propose an end-to-end cross-corpus argument mining method that uses auxiliary argument mining corpora to train models. |
| Outcome: | The proposed method outperforms models trained on a single corpus on arguments on arguments in argument mining tasks. |
Copied to clipboard
| Challenge: | Recent work in Dialogue Act (DA) classification approaches the task as a sequence labeling problem, using neural network models coupled with a Conditional Random Field (CRF) as the last layer. |
| Approach: | They propose to modify the CRF layer to take speaker-change into account and learn meaningful transition patterns conditioned on speaker-changing DA labels. |
| Outcome: | The proposed model outperforms the original model with wide margins for some DA labels. |
Copied to clipboard
| Challenge: | a recent study shows that modern neural networks understand sentences implicitly by inducing recursive structures. |
| Approach: | They propose to explicitly induce grammar by tracing the computational process of a long short-term memory network. |
| Outcome: | The proposed model can explicitly induce grammar without external knowledge . tracing the computational process of a long short-term memory network is shown to be effective . |
Copied to clipboard
| Challenge: | Existing knowledge of narrative examples is lacking and difficult to obtain. |
| Approach: | They propose a weakly supervised approach for acquiring rich temporal event knowledge across sentences in narrative stories. |
| Outcome: | The proposed approach outperforms neural network models on the narrative cloze task. |
Copied to clipboard
| Challenge: | Recent models for zero pronoun resolution in Chinese are short-sighted and do not capture semantic information for zeros and candidate antecedents. |
| Approach: | They propose to integrate a deep reinforcement learning approach to Chinese zero pronoun resolution. |
| Outcome: | The proposed approach outperforms the state-of-the-art methods in three experimental settings. |
Copied to clipboard
| Challenge: | Word2Vec embeddings have become popular representations of word meaning . similarity between two words is often assumed to be a direction-less measure, whereas relatedness is inherently directional. |
| Approach: | They propose to use word embeddings to predict asymmetric association between words from a dataset of production norms to generate thematically related words. |
| Outcome: | The proposed model predicts asymmetric association between words from a recently published dataset of production norms. |
Copied to clipboard
| Challenge: | Recent neural network models for coreference resolution are usually trained with heuristic loss functions that are computed over a sequence of local decisions. |
| Approach: | They propose an end-to-end reinforcement learning based coreference resolution model to directly optimize coreference evaluation metrics. |
| Outcome: | The proposed model achieves new state-of-the-art performance on the English OntoNotes v5.0 benchmark. |
Copied to clipboard
| Challenge: | Existing methods for semantic parsing are difficult to design and learn, especially in wideopen domains. |
| Approach: | They propose a neural semantic parsing approach which models semantic par- sing as an end-to-end semantic graph generation process. |
| Outcome: | The proposed model achieves state-of-the-art performance on Overnight dataset and gets competitive performance on Geo and Atis datasets. |
Copied to clipboard
| Challenge: | a challenge for aspect term extraction is to extract phrase-level aspect terms . a constituency lattice structure is constructed using the span annotations of constituents of a sentence . |
| Approach: | They propose to incorporate the span annotations of constituents of a sentence to leverage syntactic information in neural network models. |
| Outcome: | The proposed model outperforms existing models on two benchmark datasets. |
Copied to clipboard
| Challenge: | Existing methods for commonsense reasoning rely on human-crafted features and knowledge bases, but unsupervised learning is not feasible due to the lack of labeled training data or comprehensive knowledge bases. |
| Approach: | They propose two unsupervised models based on the Deep Structured Semantic Models framework to tackle two commonsense reasoning tasks: Winograd Schema Challenge (WSC) and Pronoun Disambiguation (PDP). |
| Outcome: | The proposed models capture contextual information in the sentence and co-reference information between pronouns and nouns, and achieve significant improvement over previous state-of-the-art approaches. |
Copied to clipboard
| Challenge: | Social media is known for its multi-cultural and multilingual interactions, a natural product of which is code-mixing. |
| Approach: | They analyze 6 million tweets produced by 27 thousand multilingual users speaking 12 other languages besides English to build predictive models to infer non-English languages users speak exclusively from their tweets. |
| Outcome: | The proposed models are based on a corpus of 6 million tweets produced by 27 thousand multilingual users speaking 12 other languages besides English . they show that content, style and syntax are the most predictive of non-English languages that users speak on Twitter. |
Copied to clipboard
| Challenge: | Neural network models have been the most successful for event detection, but they ignore syntactic relationships in the text. |
| Approach: | They propose a GRU-based model that combines syntactic information along with temporal structure through an attention mechanism. |
| Outcome: | The proposed model is competitive with existing models on a ACE2005 dataset. |
Copied to clipboard
| Challenge: | Recent studies show that shallow semantic role labeling (SRL) performance drops under out-of-domain setting. |
| Approach: | They propose to annotate a multi-domain Chinese predicate-argument dataset using a frame-free annotation methodology and strict double annotation for improving data quality. |
| Outcome: | The proposed dataset is compared with a dataset from six different domains. |
Copied to clipboard
| Challenge: | Aspect-based Sentiment Analysis (ABSA) evaluates sentiments toward specific aspects of entities within the text. |
| Approach: | They propose a method to enhance long-range dependencies between aspect and opinion words in ABSA by combining attention mechanisms with a syntax-based Graph Convolutional Network and a Mamba-Transformer module. |
| Outcome: | The proposed model outperforms state-of-the-art models on three benchmark datasets. |
Copied to clipboard
| Challenge: | generating puns with artificial intelligence techniques requires manual training and templates. |
| Approach: | They propose neural network models for homographic pun generation that can generate puns without requiring any pun data for training. |
| Outcome: | The proposed models generate homographic puns of good readability and quality without training. |
Copied to clipboard
| Challenge: | Existing work on syntactic knowledge models has not provided a clear picture of the properties required to produce proper syntaktic generalizations. |
| Approach: | They propose to evaluate syntactic knowledge of language models by varying model architectures . they find substantial differences in syntaktic generalization performance by model architecture . |
| Outcome: | The proposed model architectures outperform other architectures on a set of 34 English-language syntactic test suites. |
Copied to clipboard
| Challenge: | adequacy of a text’s characteristics with the person’s capacities and knowledge is critical in the case of . a child since her/his cognitive and linguistic skills are still under development. |
| Approach: | They propose a natural language processing task which consists in predicting the age from which a text can be understood by someone. |
| Outcome: | The proposed model outperforms psycholinguist models on a French text dataset and shows that the results are more accurate than psycholingual models. |
Copied to clipboard
| Challenge: | Existing methods to model answer selection(AS) are based on feature engineering and resource toolkits. |
| Approach: | They propose a model that captures features and information from question and answer text and repeats multiple times(hops) in a sequential fashion. |
| Outcome: | The proposed model performs on answer selection tasks and multi-level answer ranking tasks. |
Copied to clipboard
| Challenge: | Existing deep learning models for automatic readability assessment discard linguistic features traditionally used for the task. |
| Approach: | They propose to incorporate linguistic features into machine learning models by learning syntactic dense embeddings based on linguistic feature extraction. |
| Outcome: | Experiments with six data sets of two proficiency levels show that the proposed model can perform better than existing models. |
Copied to clipboard
| Challenge: | Neural network models are usually very data-hungry and performance of such models can suffer when labeled data is not available. |
| Approach: | They propose to provide models with additional analogy sources to strengthen analogy-formation . they propose to combine the analogy motivated approach with data hallucination or augmentation . |
| Outcome: | The proposed methods improve on state-of-the-art results on 46 languages, especially in low-resource settings. |
Copied to clipboard
| Challenge: | Existing grammar induction methods do not provide sufficient performance in downstream tasks. |
| Approach: | They propose an unsupervised grammar induction method for language understanding and generation using a grammar parser and a syntactic mask. |
| Outcome: | The proposed method performs better on from-scratch and pre-trained scenarios. |
Copied to clipboard
| Challenge: | In visual-grounded dialogue systems, first-person visual information about where the other speakers are and what they are paying attention to is crucial to understand their intentions. |
| Approach: | They propose a visually-grounded first-person dialogue (VFD) dataset with verbal and non-verbal responses. |
| Outcome: | The proposed dataset provides verbal and non-verbal responses for first-person visual information and recent neural network models. |
Copied to clipboard
| Challenge: | Recent approaches to fake news detection focus on textual features without external facts, which may lead to a misrepresentation of the truth. |
| Approach: | They propose a new fake news detection method that predicts the truth or false-hood of a claim based on relevant factual evidence or LLM’s inference mechanisms. |
| Outcome: | The proposed method produces the final synthesized prediction, along with well-founded facts or reasoning. |
Copied to clipboard
| Challenge: | Stronger neural network models and harder synthetic training settings are important to achieve high performance. |
| Approach: | They propose a query-based system that extracts compatible sets of events from news data . stronger neural network models and harder synthetic training settings are important to achieve high performance . |
| Outcome: | The proposed system outperforms baselines on a human-curated dataset of scenarios about real-world news topics. |
Copied to clipboard
| Challenge: | Neural networks are bringing incredible performance gains on text classification tasks, but they also require interpretability. |
| Approach: | They propose a latent model that selects a rationale and a classifier that learns from the words in the rationale alone. |
| Outcome: | The proposed model can predict expected value of penalties without REINFORCE and can be directly optimised towards a pre-specified text selection rate. |
Copied to clipboard
| Challenge: | Existing models struggle to generalize to unseen compositions of seen components . a new approach allows for disentangled representations and better generalization . |
| Approach: | They propose an extension to sequence-to-sequence models which encourage disentanglement by re-encoding source input. |
| Outcome: | The proposed extension delivers better generalization and more disentangled representations . human expressions can be understood by combining known atomic components . |
Copied to clipboard
| Challenge: | Existing approaches to discourse parsing use commonsense knowledge and linguistic constraints to integrate them into neural network models. |
| Approach: | They propose a knowledge regularization approach that integrates linguistic constraints with contexts for deriving word representations. |
| Outcome: | The proposed approach outperforms previous systems on the benchmark dataset PDTB for discourse parsing. |
Copied to clipboard
| Challenge: | Existing methods to explain neural network models are computationally inefficient for text inputs. |
| Approach: | They propose a method to implicitly detect word correlations by grouping correlated words from input text pairs together and measuring their contribution to corresponding NLP tasks. |
| Outcome: | The proposed method is evaluated with two different model architectures across four datasets. |
Copied to clipboard
| Challenge: | Technical logbooks are a challenging and under-explored text type in automated event identification. |
| Approach: | They propose a feedback strategy that resamples the training data based on its error in the prediction process. |
| Outcome: | The proposed approach provides the best results for four different neural network models trained across a suite of technical logbook datasets from distinct technical domains. |
Copied to clipboard
| Challenge: | Grapheme-to-phoneme conversion is a task of converting grapheme sequences into phoneme sequence. |
| Approach: | They propose a Thai grapheme-to-phoneme conversion method that uses neural networks to predict the similarity between a candidate and the correct pronunciation. |
| Outcome: | The proposed method can be applied to other languages than Thai . it is comparable to encoder-decoder models in accuracy and accuracy, it shows . |
Copied to clipboard
| Challenge: | Unlike Community Question Answering, where questions are mostly factoid based, forum threads are often open-ended and contain repetitive or irrelevant posts. |
| Approach: | They propose a recurrent neural network-based architecture to model the relevance of a post regarding the original post starting the thread and the novelty it brings to the discussion. |
| Outcome: | The proposed model outperforms the state-of-the-art models for text classification on different types of online forum datasets. |
Copied to clipboard
| Challenge: | Existing deep neural network models such as LSTM and tree-LSTM have a bias problem where the words in the tail of a sentence are more heavily emphasized than those in the header. |
| Approach: | They propose a capsule tree-LSTM model that uses dynamic routing to build sentence representations by assigning different weights to nodes according to their contributions to prediction. |
| Outcome: | The proposed model improves on the Stanford Sentiment Treebank and EmoBank datasets. |
Copied to clipboard
| Challenge: | Recent advances in hardware and methodology for training neural networks have enabled significant accuracy improvements across many NLP tasks. |
| Approach: | They quantify the approximate financial and environmental costs of training neural network models . they propose actionable recommendations to reduce costs and improve equity in NLP research . |
| Outcome: | The proposed recommendations address the cost and environmental costs of training neural networks for NLP. |
Copied to clipboard
| Challenge: | Table-to-text generation is a visual recognition task that uses textual descriptions from structured inputs. |
| Approach: | They propose to rethink data-to-text generation as a visual recognition task by removing the need for rendering the input in a string format. |
| Outcome: | The proposed model overcomes the challenges of linearization and input size limitations and is applicable to open-ended and controlled generation settings. |
Copied to clipboard
| Challenge: | Existing methods for word embedding compression are limited . word embeds have a considerable size and need to be compressed to deploy on edge devices . |
| Approach: | They propose a block-wise low-rank approximation method for word embedding called GroupReduce . they propose 'frequency-inverse document frequency method' and a differentiable method for weighting . |
| Outcome: | The proposed algorithm more effectively finds word weights than competitors in most cases. |
Copied to clipboard
| Challenge: | Past work has shown that many failures of systematic generalization arise from neural models’ inability to disentangle lexical phenomena from syntactic ones. |
| Approach: | They propose a lexical translation mechanism that generalizes existing copy mechanisms to incorporate learned, decontextualized, token-level translation rules. |
| Outcome: | The proposed model improves generalization on a diverse set of sequence modeling tasks drawn from cognitive science, formal semantics, and machine translation. |
Copied to clipboard
| Challenge: | Existing approaches to deep learning for open-domain dialogue include training end-to-end models to learn various conversational features like emotional content of response, symbolic transitions of dialogue contexts and persona of the agent and the user, among others. |
| Approach: | They propose a probabilistic approach using Markov Random Fields to augment existing deep-learning methods for improved next utterance prediction. |
| Outcome: | The proposed approach significantly improves the performance of existing state-of-the-art retrieval models for open-domain conversational agents. |
Copied to clipboard
| Challenge: | e-mail conversations have domain-agnostic and two-level dialogue act annotations . et al. (2017): a better understanding of asynchronous conversations. |
| Approach: | They present a large-scale corpus of e-mail conversations with domain-agnostic and two-level dialogue act annotations . they use ISO standard 24617-2 as the annotation scheme to annotate over 6,000 messages and 35,000 sentences . |
| Outcome: | The proposed model outperforms other neural networks but falls short of human performance. |
Copied to clipboard
| Challenge: | Document retrieval systems often use two styles of neural network models . dual encoder models are used for retrieval and deep re-ranking, while cross-attention models are typically used for shallow reranking. |
| Approach: | They propose a dual encoder and cross-attention neural network architectures that combine query and document representations to optimize retrieval accuracy. |
| Outcome: | The proposed architecture trades off retrieval accuracy with joint computation and offline document storage cost. |
Copied to clipboard
| Challenge: | In order for robots to collaborate with humans, it is important to share and understand their experiences through language. |
| Approach: | They propose a hierarchical Dirichlet Process-Spectral Mixture Latent Dirichlets Allocation model which learns the relationship between human motions and adverbs by capturing frequency kernels that represent motion characteristics and shared topics of a given aadverts. |
| Outcome: | The proposed model outperforms representative neural network models in terms of perplexity score and predicts more appropriate adverbs. |
Copied to clipboard
| Challenge: | Recent neural network models have achieved impressive performance on sentiment classification in English and other languages. |
| Approach: | They propose an unsupervised sentiment classification model that leverages an uncontrolled machine translation system and a language discriminator to learn a shared representation. |
| Outcome: | The proposed model outperforms other models on five language pairs. |
Copied to clipboard
| Challenge: | Recent work on latent tree learning attempts to develop models with parse-valued latent variables and train them on non-parsing tasks. |
| Approach: | They propose a model with parse-valued latent variables and a strong latent tree learning result on constituency parsing. |
| Outcome: | The proposed model outperforms all baselines and performs competitively with symbolic grammar induction systems. |
Copied to clipboard
| Challenge: | Emotion cause analysis aims to identify the reasons behind emotions . previous models focus on learning architecture with local textual information . |
| Approach: | They propose a method to extract emotion cause with hierarchical neural model and knowledge-based regularizations by sentiment lexicon and common knowledge. |
| Outcome: | The proposed method outperforms baselines on two public datasets in different languages and outperformed competitive baselines by 2.08%. |
Copied to clipboard
| Challenge: | a new study examines the role of dialogue observers in psychotherapy . the model is based on motivational interviewing, which is effective for treating addictions . |
| Approach: | They propose to model MI behavioral codes for therapists by an observer . they propose to use the observer to forecast therapist and client MI behavioral code . |
| Outcome: | The proposed model outperforms baseline models for both tasks and reveals tradeoffs in performance. |
Copied to clipboard
| Challenge: | Existing approaches to continual learning (CL) are costly and time-consuming. |
| Approach: | They propose to examine the problem of continual learning in NLP through the lens of various NLP tasks and provide a critical review of existing methods. |
| Outcome: | The proposed methods are critical to the development of CL models and provide a critical review of existing methods and datasets. |
Copied to clipboard
| Challenge: | Traditional neural network models represent word senses as vectors that are uninterpretable for humans. |
| Approach: | They propose a framework that incorporates word Sense Disambiguation (WSD) by identifying and paraphrasing ambiguous words to improve sentiment predictions. |
| Outcome: | The proposed framework improves sentiment analysis accuracy and interpretability on a downstream task without ground-truth word sense labels. |
Copied to clipboard
| Challenge: | Existing studies on keyphrase extraction neglect human reading behavior during keyphrase annotating. |
| Approach: | They propose to integrate human attention into keyphrase extraction models by an attention mechanism and combine it with neural network models. |
| Outcome: | The proposed models improve on two Twitter datasets. |
Copied to clipboard
| Challenge: | Experimental results on three news summarization datasets representative of different languages and writing styles show that our approach outperforms strong baselines by a wide margin. |
| Approach: | They propose an unsupervised approach that uses a popular ranking algorithm to compute node centrality. |
| Outcome: | The proposed approach outperforms baselines on three news summarization datasets representative of different languages and writing styles. |
Copied to clipboard
| Challenge: | Numeracy is the ability to predict the magnitude of a numeral at some specific position in a text description. |
| Approach: | They propose to use a dataset to test whether neural network models can learn numeracy . numerability is the ability to predict the magnitude of a numeral at some specific position in a text description. |
| Outcome: | The proposed task can predict the magnitude of a numeral at a specific position in a text description. |
Copied to clipboard
| Challenge: | This paper presents a recurrent neural network model to automate the analysis of students' computational thinking in problem-solving dialogue. |
| Approach: | They propose a recurrent neural network model to automate the analysis of students' computational thinking in problem-solving dialogue. |
| Outcome: | The proposed model outperforms the baseline model and outperformed the nave model by a large margin. |
Copied to clipboard
| Challenge: | Existing named entity recognition models use gazetteers to improve performance, but they are limited in coverage and do not exist in low-resource languages. |
| Approach: | They propose a method that integrates Wikipedia information into named entity models by cross-lingual entity linking. |
| Outcome: | The proposed method improves on four low-resource languages with Wikipedia . it incorporates available information from english knowledge bases into neural models . |
Copied to clipboard
| Challenge: | linguistic affiliation of languages to a common language family is traditionally carried out manually . large-scale standardized collections of multilingual wordlists and grammatical language structures could improve this . |
| Approach: | They propose to use lexical and grammatical data to classify languages into families using neural network models. |
| Outcome: | The proposed models outperform models trained on lexical and grammatical data while combining both types of data yields even better performance. |
Copied to clipboard
| Challenge: | In-context learning is an inductive bias for compositional generalization, but many deep neural architectures struggle with this ability. |
| Approach: | They propose to force a causal Transformer to in-context learn to promote compositional generalization by using earlier examples to generalize to later ones. |
| Outcome: | The proposed model can solve 'ordinary' learning problems by utilizing earlier examples to generalize to later ones, i.e., in-context learning. |
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) tasks require detecting the span and category of the entity from the text block. |
| Approach: | They propose a kNN retrieval enhancement algorithm that incorporates word segmentation information to enhance the model’s generalization ability and alleviate the problem of missing entity tokens in prediction. |
| Outcome: | The proposed method improves the performance of baseline models and achieves better or compared recognition accuracy than previous state-of-the-art models in multiple public Chinese and English datasets. |